Papers by Sabine Schulte im Walde
CCOHA: Clean Corpus of Historical American English (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to model language change in diachronic studies have been used to overcome its limitations. |
| Approach: | They propose to use the corpus of historical american english to overcome its limitations . they use a downloadable version of the corpora to remove inconsistent lemmas and malformed tokens . |
| Outcome: | The proposed corpus overcomes its main limitations without compromising its qualitative and distributional properties. |
Lexical Semantic Change Discovery (2021.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to Lexical Semantic Change Detection are limited. |
| Approach: | They propose a shift from change detection to change discovery by fine-tuning a type-based and a token-based approach on recently published German data. |
| Outcome: | The proposed models can be applied to discover new words undergoing meaning change from the full corpus vocabulary. |
Made of Steel? Learning Plausible Materials for Components in the Vehicle Repair Domain (2023.eacl-main)
Copied to clipboard
| Challenge: | a novel approach to learn domain-specific plausible materials for components in the vehicle repair domain is proposed . connecting a symptom to an underlying cause is a crucial building block for natural language understanding across domains. |
| Approach: | They propose a method to aggregate salient predictions from a set of cloze task style templates and use a Wikipedia corpus to augment the model. |
| Outcome: | The proposed approach outperforms a traditional pattern-based approach by exploiting the compositionality assumption in a cloze task style setting. |
A Systematic Search for Compound Semantics in Pretrained BERT Architectures (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing models for noun compounds have been less successful in predicting compositionality than transformers . authors: suboptimal use of encoded information may be a contributing factor . performance of transformer-based models is poor, authors say . |
| Approach: | They propose to use semantic knowledge derived from pretrained BERT to predict compositionality . they find distinct linguistic roles of heads and modifiers are reflected by differences in BERT representations . |
| Outcome: | The proposed model improves on unsupervised implementations of pretrained BERT . empirical properties such as frequency, productivity, and ambiguity affect performance . |
Combining Abstractness and Language-specific Theoretical Indicators for Detecting Non-Literal Usage of Estonian Particle Verbs (N18-4)
Copied to clipboard
| Challenge: | Existing studies on identifying nonliteral language use have focused on resource-rich languages and focused on general indicators to identify non-literal meaning. |
| Approach: | They propose to use two datasets and a random forest classifier to automatically predict literal vs. non-literal language usage for a highly frequent type of multi-word expression in a low-resource language, i.e., Estonian. |
| Outcome: | The proposed dataset outperforms a high majority baseline when combined with language-independent features of non-literal language. |
A Laypeople Study on Terminology Identification across Domains and Task Definitions (N18-2)
Copied to clipboard
| Challenge: | Existing studies on term annotation show that even experts differ in their understanding of termhood . |
| Approach: | They propose a new dataset of term annotation that examines the common understanding of what constitutes a term. |
| Outcome: | The proposed datasets show that even experts differ in their understanding of termhood . the findings suggest that there is a common understanding of what constitutes a term . |
AbsVis – Benchmarking How Humans and Vision-Language Models “See” Abstract Concepts in Images (2025.emnlp-main)
Copied to clipboard
| Challenge: | Abstract concepts like mercy and peace lack clear visual grounding, and therefore challenge humans and models to provide suitable image representations. |
| Approach: | They propose a dataset of 675 images annotated with 14,175 concept–explanation attributions from humans and two Vision-Language Models where each concept is accompanied by a textual explanation. |
| Outcome: | The proposed dataset compares human and VLM attributions in terms of diversity, abstractness, and alignment, and shows that overlapping concepts are most preferred. |
More than just Frequency? Demasking Unsupervised Hypernymy Prediction Methods (2021.findings-acl)
Copied to clipboard
| Challenge: | Using unsupervised methods of hypernymy prediction, we show that the predictions of three methods overlap and are highly correlated with frequency-based predictions. |
| Approach: | They compare unsupervised methods of hypernymy prediction to supervised methods . they show that the methods overlap and are highly correlated with frequency-based predictions . |
| Outcome: | The proposed methods overlap and are highly correlated with frequency-based predictions across English and German datasets. |
Variants of Vector Space Reductions for Predicting the Compositionality of English Noun Compounds (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to predict the degree of compositionality of noun compounds are based on comparing compounds and their constituents within a vector space and using distributional similarity as a proxy to predict their degree of semantic relatedness. |
| Approach: | They propose to use distributional similarity as a proxy to predict the semantic relatedness between the compounds and their constituents as the compound’s degree of compositionality. |
| Outcome: | The proposed methods are most successful and stable in terms of dimensionality and part-of-speech reductions. |
Projecting Embeddings for Domain Adaption: Joint Modeling of Sentiment Analysis in Diverse Domains (C18-1)
Copied to clipboard
| Challenge: | Existing domain adaptation methods for sentiment analysis are sensitive to domain differences, resulting in classifiers that perform poorly on new domains. |
| Approach: | They propose a domain adaptation problem as an embedding projection task using two mono-domain embeddable spaces and a bi-domain space to project across domains and predict sentiment. |
| Outcome: | The proposed model performs better on domains similar to state-of-the-art methods while requiring longer training times. |
VOLIMET: A Parallel Corpus of Literal and Metaphorical Verb-Object Pairs for English–German and English–French (2024.starsem-1)
Copied to clipboard
| Challenge: | Metaphorical language is a complex interplay of cultural and linguistic elements that characterizes metaphorical language . a corpus of parallel sentences containing gold standard alignments of metaphorical verb-object pairs and literal paraphrases is presented . |
| Approach: | They propose to analyze metaphorical verb-object pairs and literal paraphrases in parallel sentences from English to German and French. |
| Outcome: | The proposed corpus of 2,916 parallel sentences reveals monolingual patterns for metaphorical vs. literal uses in English . cross-lingually, the results show a rich variability in translations as well as different behaviors for the two target languages . |
Modeling Sense Structure in Word Usage Graphs with the Weighted Stochastic Block Model (2021.starsem-1)
Copied to clipboard
| Challenge: | Word Usage Graphs capture fine-grained semantic proximity distinctions between word uses. |
| Approach: | They propose to model word use Graphs using a Bayesian weighted stochastic block model and a probabilistic weightes-based model to capture fine-grained semantic proximity distinctions between word uses. |
| Outcome: | The proposed model is compared with existing models and is empirically most adequate. |
You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLP (D19-1)
Copied to clipboard
| Challenge: | Current approaches to social media modelling ignore the fact that an individual may be part of several communities which are not equally relevant in all communicative situations. |
| Approach: | They propose a model that captures the sociological phenomenon of homophily and combines it with linguistic information to make a prediction. |
| Outcome: | The proposed model significantly outperforms existing models on three different tasks and is compared with other models. |
A Domain-Specific Dataset of Difficulty Ratings for German Noun Compounds in the Domains DIY, Cooking and Automotive (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset with difficulty ratings for 1,030 closed noun compounds is presented . authors use a simple compound splitter to identify compound types in domain-specific texts . |
| Approach: | They present a German closed noun compound dataset with difficulty ratings . they used a simple compound splitter to identify compounds in texts . |
| Outcome: | The proposed dataset has difficulty ratings for 1,030 closed noun compounds extracted from domain-specific texts for do-it-ourself, cooking and automotive. |
Introducing Two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness (N18-2)
Copied to clipboard
| Challenge: | Existing datasets for low-resource language Vietnamese assess semantic similarity . a dataset for word pairs with similarity levels is needed to evaluate these models . |
| Approach: | They present two new datasets for the low-resource language Vietnamese to assess models of semantic similarity. |
| Outcome: | The two datasets are comparable to the English datasets. |
Features of Perceived Metaphoricity on the Discourse Level: Abstractness and Emotionality (2022.lrec-1)
Copied to clipboard
| Challenge: | a metaphorical discourse is more emotional and abstract than a literal one, according to a new study . a metaphorical discourse may be more abstract than literal, but it is not triggered by its emotionality or metaphoricity. |
| Approach: | They examine which features human annotators perceive as important for metaphoricity . they ask: is a metaphorical expression preceded by a more metaphorical/abstract/emotional context? |
| Outcome: | The proposed dataset shows that metaphorical discourses are more emotional and abstract than literal ones. |
Investigating Independence vs. Control: Agenda-Setting in Russian News Coverage on Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | a major challenge in the media industry has always been its targeted manipulation, says a new study . agenda-setting is a well-known phenomenon in political science . authors explore the relationship between economic indicators and mentions of foreign geopolitical entities . |
| Approach: | They investigate agenda-setting in the Russian social media landscape . they explore the relation between economic indicators and mentions of foreign geopolitical entities . |
| Outcome: | The authors examine the relationship between economic indicators and mentions of foreign geopolitical entities, as well as of Russia itself. |
To Split or Not to Split: Composing Compounds in Contextual Vector Spaces (2023.emnlp-main)
Copied to clipboard
| Challenge: | Contextual word embedding models rely on sub-word tokenization to represent single orthographic words but are often suboptimal in under-resourced contexts. |
| Approach: | They propose to use a masked language modelling task to evaluate the model's performance . they use re-trained tokenizers to pre-split compounds into constituents . |
| Outcome: | The proposed models improve on the masked language modelling task and compositionality prediction by pre-splitting compounds into constituents. |
DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches to dialectology have been limited and rarely address language variation regarding lexical meaning. |
| Approach: | They propose to use existing framework DURel and framework-embedded Word Usage Graphs to distinguish, visualize and interpret diatopic lexical semantic variation of contextualized words in Spanish from these perspectives. |
| Outcome: | The proposed dataset exploits existing frameworks for annotating word senses in context and framework-embedded Word Usage Graphs (WUGs) . it distinguishes, visualizes and interprets lexical semantic variation of contextualized words in Spanish from these two perspectives, i.e., semasiological and onomasiology. |
Concreteness vs. Abstractness: A Selectional Preference Perspective (2022.aacl-srw)
Copied to clipboard
| Challenge: | Using a collection of 5,438 nouns and 1,275 verbs, we exploit selectional preferences as a salient characteristic in classifying abstract vs. concrete words. |
| Approach: | They propose to use selectional preferences as a criterion to distinguish between concrete and abstract concepts and words. |
| Outcome: | The proposed method achieves an f1-score of 0.84 for nouns and 0.71 for verbs in classification and Spearman’s correlation of 0.86 for nonoms and 0.59% for verb. |
What Can Diachronic Contexts and Topics Tell Us about the Present-Day Compositionality of English Noun Compounds? (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to determine the semantic relatedness between compounds and constituents have applied a synchronic perspective, but this study examines what diachronic changes in contexts and semantic topics reveal about the compounds’ present-day compositionality. |
| Approach: | They propose to use two diachronic vector spaces to model compositional patterns between compounds with low and high present-day compositionality. |
| Outcome: | The proposed model performs on par with co-occurrence space and captures similar information. |
Willkommens-Merkel, Chaos-Johnson, and Tore-Klose: Modeling the Evaluative Meaning of German Personal Name Compounds (2024.lrec-main)
Copied to clipboard
Annerose Eichel, Tana Deeg, Andre Blessing, Milena Belosevic, Sabine Arndt-Lappe, Sabine Schulte im Walde
| Challenge: | Personal name compounds (PNCs) are compositions that refer to a person, such as Willkommens-Merkel ('Welcome-Meerkel') and a personal name such as Merkel. |
| Approach: | They propose to model 321 personal name compounds and their corresponding full names at discourse level and compare two approaches to assess whether a PNC is more positively or negatively evaluative . they further enrich data with personal, domain-specific, and extra-linguistic information and perform regression analyses revealing that factors including compound and modifier valence, domain, and political party membership influence how a pnc is evaluated. |
| Outcome: | The proposed model shows that the PNCs are perceived as more positively or negatively than their full name and that they are perceived to be more positive or negative. |
Bilingual Sentiment Embeddings: Joint Projection of Sentiment Across Languages (P18-1)
Copied to clipboard
| Challenge: | Existing approaches to sentiment analysis in low-resource languages lack annotated corpora or do not capture sentiment information. |
| Approach: | They propose a model that represents sentiment in a source and target language without annotated corpus. |
| Outcome: | The proposed model outperforms state-of-the-art methods on four out of six setups and captures complementary information to machine translation. |
Explaining and Improving BERT Performance on Lexical Semantic Change Detection (2021.eacl-srw)
Copied to clipboard
| Challenge: | Lexical semantic change detection is still a challenging field due to the success of type-based embeddings in SemEval-2020 Task 1 and other NLP tasks. |
| Approach: | They compare the performance of BERT embeddings with results from the word sense disambiguation dataset underlying SemEval-2020 Task 1 and the Italian follow-up task DIACR-Ita. |
| Outcome: | The proposed model outperforms token-based embeddings on lexical semantic change detection tasks. |
Predicting Degrees of Technicality in Automatic Terminology Extraction (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study has focused on term technicality, but there are still few studies on it. |
| Approach: | They semi-automatically create a German gold standard of technicality across four domains . they propose two new models to exploit general- vs. domain-specific comparisons based on vector spaces . |
| Outcome: | The proposed model outperforms previous methods in terms of general- vs. domain-specific comparisons. |
Compound or Term Features? Analyzing Salience in Predicting the Difficulty of German Noun Compounds across Domains (2021.starsem-1)
Copied to clipboard
| Challenge: | Using domain-specific vocabulary, it is important to analyse domain-related characteristics to improve the communication between lay people and experts. |
| Approach: | They focus on the interaction of compound-based lexical features (such as frequency and productivity) and terminology-based features (contrasting domain-specific and general language) across word representations and classifiers. |
| Outcome: | The proposed model shows that the interaction of compound-based lexical features and terminology-based features across word representations and classifiers is important for a broad binary distinction into ‘easy’ vs. ‘difficult’ general-language compound frequency is sufficient, but for . a more fine-grained four-class distinction it is crucial to include contrastive termhood features and compound and constituent features. |
A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains (P19-1)
Copied to clipboard
| Challenge: | Existing models for diachronic and synchronic detection of lexical semantic divergences are superficial and lack of comparison. |
| Approach: | They propose to extend benchmark models on a common state-of-the-art evaluation task . they also demonstrate that the same evaluation task and modelling approaches can be utilised for synchronic detection of domain-specific sense divergences in the field of term extraction. |
| Outcome: | The proposed model can be utilised for the detection of domain-specific sense divergences in the field of term extraction. |
Varying Vector Representations and Integrating Meaning Shifts into a PageRank Model for Automatic Term Extraction (2020.lrec-1)
Copied to clipboard
| Challenge: | a comparative study for automatic term extraction from domain-specific language using a PageRank graph algorithm with different edge-weighting methods. |
| Approach: | They propose to use a PageRank algorithm to extract automatic terms from domain-specific language using different edge-weighting methods. |
| Outcome: | The proposed model is compared with a PageRank model with different edge-weighting methods. |
Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change (N18-2)
Copied to clipboard
| Challenge: | Existing frameworks for evaluating lexical semantic change are limited . evaluation of lexicals is a major obstacle in the field of semantic change detection . |
| Approach: | They propose a framework that extends synchronic polysemy annotation to diachronic changes in lexical meaning to counteract lack of resources for evaluating computational models of lexiconal semantic change. |
| Outcome: | The proposed framework exploits an intuitive notion of semantic relatedness and distinguishes between innovative and reductive meaning changes with high inter-annotator agreement. |
Inclusive Leadership in the Age of AI: A Dataset and Comparative Study of LLMs vs. Real-Life Leaders in Workplace Action Planning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new study compares LLMs and human leaders in workplace action planning tasks . the leader success bot guides real-life leaders in generating inclusive workplace action plans . |
| Approach: | They propose a leader success bot that guides leaders in generating inclusive workplace action plans. |
| Outcome: | The Leader Success Bot guides real-life leaders in generating inclusive workplace action plans. |
Analogies in Complex Verb Meaning Shifts: the Effect of Affect in Semantic Similarity Models (N18-2)
Copied to clipboard
| Challenge: | German particle verbs are complex verb structures that combine a prefix particle with a base verb. |
| Approach: | They propose a computational model to detect and distinguish analogies in meaning shifts between German base and complex verbs using a standard similarity model. |
| Outcome: | The proposed model detects and distinguishes analogies in meaning shifts between German base and complex verbs using a standard similarity model. |